Papers with constituency parsing benchmarks
Cloze-driven Pretraining of Self-attention Networks (D19-1)
Copied to clipboard
| Challenge: | Existing work on pretraining language models has used unidirectional (left-to-right) or bi-directional (both left-to right and right-to left) LMs with loss function. |
| Approach: | They propose a bi-directional transformer model that pretrains both directions of a large language-model-inspired self-attention cloze model and propose clozing to predict each word in the training data. |
| Outcome: | The proposed model performs well on GLUE and state of the art benchmarks consistent with BERT. |